Papers with instance selection
ActiveLLM: Large Language Model-Based Active Learning for Textual Few-Shot Scenarios (2026.tacl-1)
Copied to clipboard
| Challenge: | Active learning strategies struggle with a ‘cold-start’ problem, needing substantial initial data to be effective. |
| Approach: | They propose an active learning approach that leverages Large Language Models such as GPT-4, o1, Llama 3, or Mistral Large for selecting instances. |
| Outcome: | The proposed approach outperforms existing methods ADAPET, PERFECT, and SetFit in few-shot scenarios and can be extended to non-few scenarios. |
Distant Supervision from Disparate Sources for Low-Resource Part-of-Speech Tagging (D18-1)
Copied to clipboard
| Challenge: | Low-resource languages lack manual annotated data to learn basic models such as part-of-speech (POS) taggers. |
| Approach: | They propose a cross-lingual neural part-of-speech tagger that learns from disparate sources of distant supervision in a uniform framework. |
| Outcome: | The proposed model scales to hundreds of low-resource languages without access to gold annotated data. |
Noisy Multi-Label Text Classification via Instance-Label Pair Correction (2024.findings-naacl)
Copied to clipboard
| Challenge: | Noise is a significant challenge for machine learning models, especially deep learning models. |
| Approach: | They propose a holistic selection metric that identifies noisy pairs while considering global loss information and instance-specific ranking information. |
| Outcome: | The proposed approach significantly improves performance in noisy multi-label text classification tasks. |
Posterior-regularized REINFORCE for Instance Selection in Distant Supervision (N19-1)
Copied to clipboard
| Challenge: | Existing methods to train unbiased methods such as REINFORCE take time to train. |
| Approach: | They propose to use posterior regularization to integrate domain-specific rules in instance selection using REINFORCE to improve the performance of the relation classifier trained on cleaned distant supervision datasets. |
| Outcome: | The proposed method improves the performance of the relation classifier trained on cleaned distant supervision dataset as well as the efficiency of the REINFORCE training. |